Picture for Junda Wu

Junda Wu

SkillBrew: Multi-Objective Curation of Skill Banks for LLM Agents

Add code
May 28, 2026
Viaarxiv icon

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking

Add code
May 13, 2026
Viaarxiv icon

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization

Add code
May 11, 2026
Viaarxiv icon

FERA: Uncertainty-Aware Federated Reasoning for Large Language Models

Add code
May 11, 2026
Viaarxiv icon

Skill-R1: Agent Skill Evolution via Reinforcement Learning

Add code
May 10, 2026
Viaarxiv icon

A Survey on LLM-based Conversational User Simulation

Add code
Apr 27, 2026
Viaarxiv icon

WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning

Add code
Feb 19, 2026
Viaarxiv icon

AMPS: Adaptive Modality Preference Steering via Functional Entropy

Add code
Feb 13, 2026
Viaarxiv icon

Evaluation on Entity Matching in Recommender Systems

Add code
Jan 23, 2026
Viaarxiv icon

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

Add code
Jan 09, 2026
Viaarxiv icon